Bengaluru: Humyn Labs, a Physical AI research lab focused on accelerating the deployment-readiness of robots in the real world, has launched the second edition of BRIDGE, a benchmark report examining the gap between human speech and voice AI models.
The report assesses how effectively voice AI models can support human-robot interaction and task execution in real-world environments.
The Humyn Labs BRIDGE benchmark evaluated 23 voice AI models across 23 languages using real-world noisy conversations. The benchmark compares leading models, including Sarvam v3, Gemini 3 Pro and ElevenLabs, across languages, accents and dialects.
The Humyn Labs report examines voice models using seven core metrics, including overlapping speech, conversational density and code-switching, alongside other conditions designed to mirror real-world environments.
The second edition of BRIDGE covers Indic languages as well as Latin American Spanish, Brazilian Portuguese and Vietnamese.
The benchmark is based on more than 200 hours of human-verified real-world audio collected across two to three districts for each language.
With 5.5 billion people speaking languages other than English, the ability of AI systems to understand speech across different linguistic and conversational contexts is an increasingly important measure of AI reliability and inclusivity, according to the Humyn Labs benchmark.
BRIDGE identifies impact of overlapping speech and dialects
The Humyn Labs findings show that overlapping speech significantly affects voice AI accuracy. The average error rate increased from 41.2% to 45.2% when overlapping speech was present.
Dialect differences also produced significant variations in performance. Bengali recorded a 42.4% error rate in its standard form, compared with 51.0% in a regional dialect outside Kolkata. The difference of nearly nine percentage points was attributed to dialect alone.
The Humyn Labs benchmark found that the dialect gap extends beyond Indic languages. Spanish showed a similar pattern, with Argentinian Spanish recording a 7.85% word error rate compared with 16.04% for Venezuelan Spanish. The error rate therefore more than doubled between dialects of the same language.
Voice AI Models: Selection Affects Real-World Voice AI Accuracy
The Humyn Labs report also found substantial differences between individual voice AI models. ElevenLabs was the top performer across five non-Indic languages, recording an average error rate of 5.8%.
By comparison, GPT-4o-mini-transcribe recorded a 24.6% error rate on the same audio. The Humyn Labs benchmark noted that this represented a difference of more than four times between the two models, highlighting the impact of provider selection on real-world transcription accuracy.
The benchmark also examined the effect of long pauses on voice AI performance. In Brazilian Portuguese calls, conversations with gaps exceeding 150 seconds recorded an 18.8% error rate, compared with 12.4% for shorter gaps of around 35 seconds.
According to the Humyn Labs report, the finding indicates that models losing conversational context after extended pauses is not an issue limited to Indic languages.
Also Read: VYOMA Innovation Challenge Finale Showcases 10 Open-Source AI Prototypes
Voice AI Models: Humyn Labs Highlights Business Implications of Voice Accuracy
Manish Agarwal, Co-Founder, Humyn Labs, stated, “Voice is a critical interface for Physical AI, and therefore voice accuracy becomes a business imperative, not just a technical metric.
If Voice AI models cannot understand overlapping speech, interruptions, code-switching, pauses and the diversity of languages people use every day, that gap ultimately impacts customer experience, automation and trust.
BRIDGE is designed to help Physical AI & Voice AI builders and enterprises evaluate models against the complexity of real-world conversations and understand whether they are truly ready to scale across markets.”
The Humyn Labs BRIDGE benchmark also separates script choice from genuine transcription errors, addressing a distinction that standard scoring can miss.
In Bengali, Soniox and Sarvam v3 recorded raw error rates of approximately 20% to 21%. Around half of these errors resulted from English loanwords being written in a different script rather than from the models mishearing the words.
Gemini 3 Pro recorded a 9.2% error rate, with only 1.6 percentage points attributed to script mismatch, according to the Humyn Labs findings.
Different Voice AI Models Fail in Different Ways
The Humyn Labs benchmark found that models can produce similar overall error rates while failing in structurally different ways.
According to the report, 19 of the 23 models tested misheard words and substituted incorrect words. A smaller group, including OpenAI’s transcribe models, Speechmatics and Gnani Vachana, was found to fail primarily through omission.
For these models, deletions accounted for 38–39% of their total errors, the Humyn Labs report said.
Gemini Flash displayed another failure pattern through fabrication. The model introduced words that were not spoken, adding invented content equivalent to 9.3% of the reference transcript’s length.
The Humyn Labs report highlighted the commercial importance of distinguishing between these failure modes, noting that a workflow capable of absorbing a missing word may not be able to tolerate fabricated content.
Voice AI Models: BRIDGE Evaluates Economics of Model Routing
The Humyn Labs benchmark also examined the economics of routing audio between different voice AI models.
The best single model achieved a 10.7% loanword-adjusted error rate, while selecting the best model for each language reduced the rate to 9.7%. A theoretical best-model-per-call approach achieved an 8.9% error rate.
One model won 78.6% of the files outright, according to the Humyn Labs report.
Ishank Gupta, Co-Founder, Humyn Labs, adds, “Physical AI cannot learn the real world through vision alone. Sound carries information about people, actions, distance, environment and intent and for a robot operating alongside humans, being able to interpret that signal reliably is fundamental.
BRIDGE provides the evaluation layer that has been missing for this modality: testing speech models not just on words, but across conditions and context that mirror real-world environments.”
The second edition of Humyn Labs BRIDGE is positioned as an evaluation benchmark for Physical AI and Voice AI builders and enterprises seeking to assess speech models under the complexity of real-world conversations, languages, accents, dialects and conversational conditions.







